Tag
2 articles
This article explains tile-based GPU programming concepts, focusing on NVIDIA's cuTile and Triton frameworks, and how they enable efficient Flash Attention in large language models.
Learn how to work with AI chip architectures similar to Alibaba's Zhenwu M890 by setting up your development environment, creating neural networks, and optimizing performance using Python frameworks like TensorFlow and PyTorch.